Papers with explainable medical QA
Benchmarking Large Language Models on Answering and Explaining Challenging Medical Questions (2025.naacl-long)
Copied to clipboard
| Challenge: | Medical board exams or general clinical questions do not capture the complexity of real clinical cases. |
| Approach: | They construct two datasets that are structured as multiple-choice question-answering tasks accompanied by expert-written explanations. |
| Outcome: | The proposed datasets are harder than previous benchmarks. |